【文章标题】:[AINews] 10% worse, 100x cheaper, 10000x faster: Why Simulation is taking over

[AINews] 差10%,便宜100倍,快10000倍:为什么模拟正在接管

【文章正文】:

By AI standards today is a pretty quiet Friday, so it’s time to take a step back and reflect on what is really going on. If you read our

2025 reading list

, and followed our coverage of

Z.ai GLM

, understood

the Poolside pivot

, been following our

AI for Science themes

, and tuned in to today’s

Simile pod

, you not only are one of the biggest readers of Latent Space, you will probably also arrive at this mental model:

按照AI的标准,今天是一个相当安静的星期五,所以是时候退后一步,反思一下真正发生了什么。如果你读过我们的

2025 阅读清单

,关注过我们对

Z.ai GLM

的报道,理解了

Poolside 的转型

,一直在关注我们的

AI for Science 主题

,并收听了今天的

Simile 播客

,那么你不仅是 Latent Space 的最大读者之一,你很可能也会得出这样的心智模型:

Every year since 2022, one more component of the pipeline that produces machine intelligence has flipped from human-made to model-made. Not gradually, and not evenly — each flip has a patient zero, a paper or product where the synthetic version first became load-bearing at a frontier lab, and from there on, the future is simply here but not yet productionized.

自2022年以来,每一年,产生机器智能的流水线中都会有一个组件从人造翻转为模型造。不是逐渐发生的,也不是均匀发生的——每一次翻转都有一个零号病人,一篇论文或一个产品,其中合成版本首次在前沿实验室中成为承重部件,从那时起,未来就已经到来,只是尚未产品化。

And if you squint, what we used to call “synthetic data” and “synthetic rubrics” and “AI researcher” and “end to end RL environments” is just

increasingly ambitious human simulation

  • 10% worse, but 100x cheaper and 10,000x faster.

如果你眯起眼睛看,我们过去所说的“合成数据”、“合成评分标准”、“AI研究员”和“端到端RL环境”只不过是

越来越雄心勃勃的人类模拟

——差10%,但便宜100倍,快10000倍。

Stage 1: The reward signal (2022)

阶段1:奖励信号(2022年)

The first thing to go synthetic was, counterintuitively, the judge.

InstructGPT

established the now-canonical trick: collect human preferences once, train a

reward model

, and let the policy optimize against the model rather than the humans. From the policy’s point of view, the thing dispensing approval was already an LLM.

Constitutional AI

pushed further and had the AI critique itself against a set of principles (RLAIF), and

Lee et al.

later showed AI feedback matching human feedback at a fraction of the cost. By the time

LLM-as-judge

became the default eval methodology (MT-Bench, AlpacaEval), the entire approval apparatus — reward, critique, evaluation — ran on models judging models.

第一个变成合成物的,出乎意料地,是裁判。

InstructGPT

确立了如今已成为经典的做法:收集一次人类偏好,训练一个

奖励模型

,让策略针对模型而不是人类进行优化。从策略的角度来看,发放认可的那个东西已经是一个LLM。

Constitutional AI

更进一步,让AI根据一组原则进行自我批判(RLAIF),而

Lee等人

后来表明,AI反馈以极低的成本与人类反馈相匹配。等到

LLM-as-judge

成为默认评估方法(MT-Bench、AlpacaEval)时,整个认可装置——奖励、批判、评估——都运行在模型评判模型之上。

Stage 2: The training data (2023)

阶段2:训练数据(2023年)

Microsoft’s Phi series made the argument in its title:

Textbooks Are All You Need

. A small model trained on LLM-synthesized, textbook-quality data punched far above its parameter count, and

phi-1.5

confirmed it wasn’t a fluke. Apple’s

WRAP

generalized the move: don’t just generate data,

rephrase the entire web

with an LLM, and pretraining gets roughly 3x more efficient. From there the pipeline industrialized — NVIDIA’s

Nemotron-4 340B

shipped with a permissively licensed synthetic data generation pipeline as a headline feature, and by 2025 reasoning-trace corpora (chains of thought generated by strong reasoners) had become a standard pretraining and mid-training ingredient. The corpus, the thing that was supposed to be the irreducibly human input, was now substantially model-written.

微软的Phi系列在标题中提出了这个论点:

教科书就是你所需要的全部

。一个在LLM合成的、教科书质量的数据上训练的小模型,其表现远远超出了它的参数数量,而

phi-1.5

证实了这不是侥幸。苹果的

WRAP

将这一做法推广:不要只是生成数据,

用LLM重新表述整个网络

,预训练的效率大约提高了3倍。从那时起,这条流水线实现了工业化——NVIDIA的

Nemotron-4 340B

发布时,将一条许可宽松的合成数据生成流水线作为头条功能,而到2025年,推理轨迹语料库(由强推理器生成的思维链)已成为标准的预训练和中期训练原料。语料库,这个本应是不可还原的人类输入的东西,现在基本上是由模型编写的。

Stage 3: The teacher (2023)

阶段3:教师(2023年)

Weeks after ChatGPT’s API opened, Stanford’s

Alpaca

demonstrated that a $600 fine-tune on GPT-generated instructions could clone much of a frontier model’s behavior.

Vicuna

did it with shared conversations;

Orca

did it with rich teacher explanations rather than bare answers. The technique matured from imitation into a proper training discipline —

on-policy generalized knowledge distillation

fixed the train/inference mismatch — and reached its cultural peak when

DeepSeek-R1

shipped a family of distilled models alongside the flagship, making “the teacher is a model” the default assumption for every small model release since.

在ChatGPT的API开放几周后,斯坦福的

Alpaca

证明,对GPT生成的指令进行600美元的微调,就能克隆出前沿模型的大部分行为。

Vicuna

用共享对话做到了这一点;

Orca

则用丰富的教师解释而不是光秃秃的答案做到了这一点。这项技术从模仿成熟为一门正式的训练学科——

同策略广义知识蒸馏

修复了训练/推理不匹配的问题——并在

DeepSeek-R1

随旗舰模型一起发布一系列蒸馏模型时达到了文化顶峰,使得“教师是一个模型”成为此后每一次小模型发布的默认假设。

Stage 4: The curriculum (2024)

阶段4:课程(2024年)

Stages 1–3 made the inputs synthetic; stage 4 is where the loop starts closing on itself, because the model begins deciding

what to learn next

. The pieces existed early —

Self-Instruct

(models writing their own instruction sets) and

STaR

(models bootstrapping their own reasoning traces) are both 2022 — but the flip came when Meta’s

Self-Rewarding Language Models

and

SPIN

showed a model could generate its own tasks, judge its own outputs, and improve past the ceiling of its human preference data. Curriculum design — historically the most artisanal part of ML, the taste-driven choice of what to train on next — became something models do to themselves.

阶段1至3使输入变成了合成物;阶段4是循环开始自我闭合的地方,因为模型开始决定

接下来学习什么

。这些拼图很早就存在——

Self-Instruct

(模型编写自己的指令集)和

STaR

(模型引导自己的推理轨迹)都是2022年的——但翻转发生在Meta的

自我奖励语言模型

和

SPIN

表明一个模型可以生成自己的任务、评判自己的输出,并改进到超越其人类偏好数据的天花板时。课程设计——历史上是机器学习中最具手工艺色彩的部分,是基于品味选择接下来训练什么——变成了模型对自己做的事情。

Stage 5: The researcher (2026)

阶段5:研究员(2026年)

The assistance era (Copilot, then SWE-agents) kept a human choosing the experiments. The discovery era did not. DeepMind’s

AlphaEvolve

evolved genuinely new algorithms in 2025, and Sakana’s

AI Scientist

(now in

Nature

!) sketched the full paper-writing pipeline. The big moment was Karpathy’s

autoresearch

in March 2026: a deliberately minimal ratchet loop where a coding agent modifies a real LLM training setup, runs a five-minute experiment, keeps the change only if validation loss improves, and repeats overnight. His own extended run stacked 700 experiments into 20 kept improvements, cutting time-to-GPT-2 from 2.02 to 1.80 hours — real, transferable code changes found while he slept.

辅助时代(Copilot,然后是SWE智能体)让人类来选择实验。发现时代则不然。DeepMind的

AlphaEvolve

在2025年进化出了真正的新算法,而Sakana的

AI Scientist

(如今已登上

Nature

!)勾勒出了完整的论文写作流水线。重大时刻是Karpathy在2026年3月的

autoresearch

:一个刻意极简的棘轮循环,其中编码智能体修改真实的LLM训练设置,运行五分钟实验,只有在验证损失改善时才保留更改,然后整夜重复。他自己的一次延长运行将700个实验堆叠成20个被保留的改进,将训练到GPT-2的时间从2.02小时缩短到1.80小时——在他睡觉时发现的真实、可迁移的代码更改。

Stage 6: The environment (2026)

阶段6:环境(2026年)

RL’s scaling bottleneck moved from the model to the environment: you need thousands of executable, verifiable, professionally realistic task worlds, a

RL的扩展瓶颈从模型转移到了环境:你需要成千上万个可执行、可验证、专业上逼真的任务世界,一个